Improving the calculation of statistical significance in genome-wide scans.
نویسندگان
چکیده
Calculations of the significance of results from linkage analysis can be performed by simulation or by theoretical approximation, with or without the assumption of perfect marker information. Here we concentrate on theoretical approximation. Our starting point is the asymptotic approximation formula presented by Lander and Kruglyak (1995, Nature Genetics, 11, 241--247), incorporating the effect of finite marker spacing as suggested by Feingold et al. (1993, American Journal of Human Genetics, 53, 234--251). We consider two distinct ways in which this formula can be improved. Firstly, we present a formula for calculating the crossover rate rho for a pedigree of general structure. For a pedigree set, these values may then be weighted into an overall crossover rate which can be used as input to the original approximation formula. Secondly, the unadjusted p -value formula is based on the assumption of a Normally distributed nonparametric linkage (NPL) score. This leads to conservative or anti-conservative p -values of varying magnitude depending on the pedigree set structure. We adjust for non-Normality by calculating the marginal distribution of the NPL score under the null hypothesis of no linkage with an arbitrarily small error. The NPL score is then transformed to have a marginal standard Normal distribution and the transformed maximal NPL score, together with a slightly corrected value of the overall crossover rate, is inserted into the original formula in order to calculate the p -value. We use pedigrees of seven different structures to compare the performance of our suggested approximation formula to the original approximation formula, with and without skewness correction, and to results found by simulation. We also apply the suggested formula to two real pedigree set structure examples. Our method generally seems to provide improved behavior, especially for pedigree sets which show clear departure from Normality, in relation to the competing approximations.
منابع مشابه
Statistical Significance for Genome-Wide Experiments
With the increase in genome-wide experiments and the sequencing of multiple genomes, the analysis of large data sets has become commonplace in biology. It is often the case that thousands of features in a genome-wide data set are tested against some null hypothesis, where many features are expected to be significant. Here we propose an approach to statistical significance in the analysis of gen...
متن کاملEstimating odds ratios in genome scans: an approximate conditional likelihood approach.
In modern whole-genome scans, the use of stringent thresholds to control the genome-wide testing error distorts the estimation process, producing estimated effect sizes that may be on average far greater in magnitude than the true effect sizes. We introduce a method, based on the estimate of genetic effect and its standard error as reported by standard statistical software, to correct for this ...
متن کاملAssessing genome-wide statistical significance for large p small n problems.
Assessing genome-wide statistical significance is an important issue in genetic studies. We describe a new resampling approach for determining the appropriate thresholds for statistical significance. Our simulation results demonstrate that the proposed approach accurately controls the genome-wide type I error rate even under the large p small n situations.
متن کاملOn Assessing Genome-Wide Statistical Significance for Large p Small n Problems
Assessing genome-wide statistical significance is an important issue in genetic studies. We describe a new resampling approach for determining the appropriate thresholds for statistical significance. Our simulation results demonstrate that the proposed approach accurately controls the genome-wide type I error rate even under the large p small n situations.
متن کاملGenome Wide Association Studies, Next Generation Sequencing and Their Application in Animal Breeding and Genetics: A Review
Recently genetic studies have been revolutionized by next generation sequencing (NGS) technology, and it is expected that the use of this technology will largely eliminate defects in the methods of association studies. The NGS technology is becoming the premier tool in genetics. However, at the moment the use of this method is limited especially in the livestock due to high cost and computation...
متن کاملP30: Are There Anxious Genes?
Anxiety comprises many clinical descriptions and phenotypes. A genetic predisposition to anxiety is undoubted; however, the nature and extent of that contribution is still unclear. Extensive genetic studies of the serotonin (5-hydroxytryptamine, 5-HT) transporter (5-HTT) gene have revealed how variation in gene expression can be correlated with anxiety phenotypes. Complete genome-wide linkage s...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید
ثبت ناماگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید
ورودعنوان ژورنال:
- Biostatistics
دوره 6 4 شماره
صفحات -
تاریخ انتشار 2005